<!DOCTYPE html>
<html class="client-nojs vector-feature-night-mode-disabled vector-feature-language-in-header-enabled vector-feature-language-in-main-page-header-disabled vector-feature-page-tools-pinned-disabled vector-feature-toc-pinned-clientpref-1 vector-feature-main-menu-pinned-disabled vector-feature-limited-width-clientpref-1 vector-feature-limited-width-content-enabled vector-feature-custom-font-size-clientpref-1 vector-feature-appearance-pinned-clientpref-1 vector-sticky-header-enabled" lang="en" dir="ltr"><head>
<meta charset="UTF-8">
<title>Text segmentation</title>
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<link rel="canonical" href="https://en.wikipedia.org/wiki/Text_segmentation"> <link href="./mw/ext.cite.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.icons.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.search.codex.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/user.styles.css" rel="stylesheet" type="text/css">
<meta name="ResourceLoaderDynamicStyles" content="">
<link rel="stylesheet" type="text/css" href="./mw/site.styles.css">
<link rel="stylesheet" type="text/css" href="./mw/noscript.css">
<link rel="stylesheet" type="text/css" href="./footer.css">
<link rel="stylesheet" type="text/css" href="./vector-2022.css">
</head>
<body class="skin--responsive skin-vector skin-vector-search-vue mediawiki ltr sitedir-ltr mw-hide-empty-elt ns-0 ns-subject page-Text_segmentation rootpage-Text_segmentation skin-vector-2022 action-view">
<div class="mw-page-container">
<div class="mw-page-container-inner">
<div class="mw-content-container">
<main id="content" class="mw-body">
<header class="mw-body-header vector-page-titlebar">
<h1 id="firstHeading" class="firstHeading mw-first-heading">
<span id="openzim-page-title" class="mw-page-title-main"><span class="mw-page-title-main">Text segmentation</span></span>
</h1>
</header>
<a id="top"></a>
<div id="bodyContent" class="vector-body ve-init-mw-desktopArticleTarget-targetContainer" aria-labelledby="firstHeading" data-mw-ve-target-container="">
<div id="mw-content-text" class="mw-body-content mw-content-ltr" lang="en" dir="ltr"><div class="mw-content-ltr mw-parser-output" lang="en" dir="ltr">
<p class="mw-empty-elt">
</p>
<style data-mw-deduplicate="TemplateStyles:r1305433154">
/* start https://en.wikipedia.org/ */
.mw-parser-output .ambox{border:1px solid #a2a9b1;border-left:10px solid #36c;background-color:#fbfbfb;box-sizing:border-box}.mw-parser-output .ambox+link+.ambox,.mw-parser-output .ambox+link+style+.ambox,.mw-parser-output .ambox+link+link+.ambox,.mw-parser-output .ambox+.mw-empty-elt+link+.ambox,.mw-parser-output .ambox+.mw-empty-elt+link+style+.ambox,.mw-parser-output .ambox+.mw-empty-elt+link+link+.ambox{margin-top:-1px}html body.mediawiki .mw-parser-output .ambox.mbox-small-left{margin:4px 1em 4px 0;overflow:hidden;width:238px;border-collapse:collapse;font-size:88%;line-height:1.25em}.mw-parser-output .ambox-speedy{border-left:10px solid #b32424;background-color:#fee7e6}.mw-parser-output .ambox-delete{border-left:10px solid #b32424}.mw-parser-output .ambox-content{border-left:10px solid #f28500}.mw-parser-output .ambox-style{border-left:10px solid #fc3}.mw-parser-output .ambox-move{border-left:10px solid #9932cc}.mw-parser-output .ambox-protection{border-left:10px solid #a2a9b1}.mw-parser-output .ambox .mbox-text{border:none;padding:0.25em 0.5em;width:100%}.mw-parser-output .ambox .mbox-image{border:none;padding:2px 0 2px 0.5em;text-align:center}.mw-parser-output .ambox .mbox-imageright{border:none;padding:2px 0.5em 2px 0;text-align:center}.mw-parser-output .ambox .mbox-empty-cell{border:none;padding:0;width:1px}.mw-parser-output .ambox .mbox-image-div{width:52px}@media(min-width:720px){.mw-parser-output .ambox{margin:0 10%}}@media print{body.ns-0 .mw-parser-output .ambox{display:none!important}}
/* end https://en.wikipedia.org/ */
</style>
<p><b>Text segmentation</b> is the process of dividing written text into meaningful units, such as words, <a href="Sentence_(linguistics)" title="Sentence (linguistics)">sentences</a>, or <a href="Topic_(linguistics)" class="mw-redirect" title="Topic (linguistics)">topics</a>. The term applies both to <a href="Mental_process" class="mw-redirect" title="Mental process">mental processes</a> used by humans when reading text, and to artificial processes implemented in computers, which are the subject of <a href="Natural_language_processing" title="Natural language processing">natural language processing</a>. The problem is non-trivial, because while some written languages have explicit word boundary markers, such as the word spaces of written English and the distinctive initial, medial and final letter shapes of <a href="Arabic_language" class="mw-redirect" title="Arabic language">Arabic</a>, such signals are sometimes ambiguous and not present in all written languages.
</p><p>Compare <a href="Speech_segmentation" title="Speech segmentation">speech segmentation</a>, the process of dividing speech into linguistically meaningful portions.
</p>
<meta property="mw:PageProp/toc">
<div class="mw-heading mw-heading2"><h2 id="Segmentation_problems">Segmentation problems</h2></div>
<div class="mw-heading mw-heading3"><h3 id="Word_segmentation">Word segmentation</h3></div>
<style data-mw-deduplicate="TemplateStyles:r1236090951">
/* start https://en.wikipedia.org/ */
.mw-parser-output .hatnote{font-style:italic}.mw-parser-output div.hatnote{padding-left:1.6em;margin-bottom:0.5em}.mw-parser-output .hatnote i{font-style:normal}.mw-parser-output .hatnote+link+.hatnote{margin-top:-0.5em}@media print{body.ns-0 .mw-parser-output .hatnote{display:none!important}}
/* end https://en.wikipedia.org/ */
</style><div role="note" class="hatnote navigation-not-searchable">See also: <a href="Word#Word_boundaries" title="Word">Word § Word boundaries</a></div>
<p>Word segmentation is the problem of dividing a string of written language into its component words.
</p><p>In English and many other languages using some form of the <a href="Latin_alphabet" title="Latin alphabet">Latin alphabet</a>, the <a href="Space_(punctuation)" title="Space (punctuation)">space</a> is a good approximation of a <a href="Word_divider" title="Word divider">word divider</a> (word <a href="Delimiter" title="Delimiter">delimiter</a>), although this concept has limits because of the variability with which languages <a href="Emic_and_etic" title="Emic and etic">emically</a> regard <a href="Collocation" title="Collocation">collocations</a> and <a href="Compound_(linguistics)" title="Compound (linguistics)">compounds</a>. Many <a href="English_compound#Compound_nouns" title="English compound">English compound nouns</a> are variably written (for example, <i><a href="Icebox" title="Icebox">ice box = ice-box = icebox</a></i>; <i><a href="Sty" title="Sty">pig sty = pig-sty = pigsty</a></i>) with a corresponding variation in whether speakers think of them as <a href="Noun_phrase" title="Noun phrase">noun phrases</a> or single nouns; there are trends in how norms are set, such as that open compounds often tend eventually to solidify by widespread convention, but variation remains systemic. In contrast, <a href="German_nouns#Compounds" title="German nouns">German compound nouns</a> show less orthographic variation, with solidification being a stronger norm.
</p><p>However, the equivalent to the word space character is not found in all written scripts, and without it word segmentation is a difficult problem. Languages which do not have a trivial word segmentation process include Chinese, Japanese, where <a href="Sentences" title="Sentences">sentences</a> but not words are delimited, <a href="Thai_language" title="Thai language">Thai</a> and <a href="Lao_language" title="Lao language">Lao</a>, where phrases and sentences but not words are delimited, and <a href="Vietnamese_language" title="Vietnamese language">Vietnamese</a>, where syllables but not words are delimited.
</p><p>In some writing systems however, such as the <a href="Ge'ez_script" class="mw-redirect" title="Ge'ez script">Ge'ez script</a> used for <a href="Amharic" title="Amharic">Amharic</a> and <a href="Tigrinya_language" title="Tigrinya language">Tigrinya</a> among other languages, words are explicitly delimited (at least historically) with a non-whitespace character.
</p><p>The <a href="Unicode_Consortium" title="Unicode Consortium">Unicode Consortium</a> has published a <i>Standard Annex on Text Segmentation</i>,<sup id="cite_ref-1" class="reference"><a href="#cite_note-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup> exploring the issues of segmentation in multiscript texts.
</p><p><b>Word splitting</b> is the process of <a href="Parsing" title="Parsing">parsing</a> <a href="Concatenated" class="mw-redirect" title="Concatenated">concatenated</a> text (i.e. text that contains no spaces or other word separators) to infer where word breaks exist.
</p><p>Word splitting may also refer to the process of <a href="Syllabification" title="Syllabification">hyphenation</a>.
</p><p>Some scholars have suggested that modern Chinese should be written in word segmentation, with
spaces between words like written English.<sup id="cite_ref-2" class="reference"><a href="#cite_note-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup> Because there are ambiguous texts where only the author knows the intended meaning. For example, "美国会不同意。" may mean "美国 会 不同意。" (The US will not agree.) or "美 国会 不同意。" (The US Congress does not agree). For more details, see <a href="Chinese_word-segmented_writing" title="Chinese word-segmented writing">Chinese word-segmented writing</a>.
</p>
<div class="mw-heading mw-heading3"><h3 id="Intent_segmentation">Intent segmentation</h3></div>
<p>Intent segmentation is the problem of dividing written words into keyphrases (2 or more group of words).
</p><p>In English and all other languages the core intent or desire is identified and become the corner-stone of the keyphrase Intent segmentation. Core product/service, idea, action & or thought anchor the keyphrase.
</p><p>"[All things are made of <b>atoms</b>]. [Little <b>particles</b> that move] [around in perpetual <b>motion</b>], [attracting each <b>other</b>] [when they are a little <b>distance</b> apart], [but <b>repelling</b>] [upon being <b>squeezed</b>] [into <b>one another</b>]."
</p>
<div class="mw-heading mw-heading3"><h3 id="Sentence_segmentation">Sentence segmentation</h3></div>
<div role="note" class="hatnote navigation-not-searchable">See also: <a href="Sentence_boundary_disambiguation" title="Sentence boundary disambiguation">Sentence boundary disambiguation</a></div>
<p>Sentence segmentation is the problem of dividing a string of written language into its component <a href="Sentence_(linguistics)" title="Sentence (linguistics)">sentences</a>. In English and some other languages, using punctuation, particularly the <a href="Full_stop" title="Full stop">full stop</a>/period character is a reasonable approximation. However even in English this problem is not trivial due to the use of the full stop character for abbreviations, which may or may not also terminate a sentence. For example, <i>Mr.</i> is not its own sentence in "<i>Mr. Smith went to the shops in Jones Street."</i> When processing plain text, tables of abbreviations that contain periods can help prevent incorrect assignment of sentence boundaries.
</p><p>As with word segmentation, not all written languages contain punctuation characters that are useful for approximating sentence boundaries.
</p>
<div class="mw-heading mw-heading3"><h3 id="Topic_segmentation">Topic segmentation</h3></div>
<div role="note" class="hatnote navigation-not-searchable">See also: <a href="Document_classification" title="Document classification">Document classification</a></div>
<p>Topic analysis consists of two main tasks: topic identification and text segmentation. While the first is a simple <a href="Machine_learning" title="Machine learning">classification</a> of a specific text, the latter case implies that a document may contain multiple topics, and the task of computerized text segmentation may be to discover these topics automatically and segment the text accordingly. The topic boundaries may be apparent from section titles and paragraphs. In other cases, one needs to use techniques similar to those used in <a href="Document_classification" title="Document classification">document classification</a>.
</p><p>Segmenting the text into <a href="Topic_(linguistics)" class="mw-redirect" title="Topic (linguistics)">topics</a> or <a href="Discourse" title="Discourse">discourse</a> turns might be useful in some natural processing tasks: it can improve <a href="Information_retrieval" title="Information retrieval">information retrieval</a> or <a href="Speech_recognition" title="Speech recognition">speech recognition</a> significantly (by indexing/recognizing documents more precisely or by giving the specific part of a document corresponding to the query as a result). It is also needed in <a href="Topic_detection" class="mw-redirect" title="Topic detection">topic detection</a> and tracking systems and <a href="Text_summarization" class="mw-redirect" title="Text summarization">text summarizing</a> problems.
</p><p>Many different approaches have been tried:<sup id="cite_ref-3" class="reference"><a href="#cite_note-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-4" class="reference"><a href="#cite_note-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup> e.g. <a href="Hidden_Markov_model" title="Hidden Markov model">HMM</a>, <a href="Lexical_chain" title="Lexical chain">lexical chains</a>, passage similarity using word <a href="Co-occurrence" title="Co-occurrence">co-occurrence</a>, <a href="Cluster_analysis" title="Cluster analysis">clustering</a>, <a href="Topic_modeling" class="mw-redirect" title="Topic modeling">topic modeling</a>, etc.
</p><p>It is quite an ambiguous task – people evaluating the text segmentation systems often differ in topic boundaries. Hence, text segment evaluation is also a challenging problem.
</p>
<div class="mw-heading mw-heading3"><h3 id="Other_segmentation_problems">Other segmentation problems</h3></div>
<p>Processes may be required to segment text into segments besides mentioned, including <a href="Morpheme" title="Morpheme">morphemes</a> (a task usually called <a href="Morphology_(linguistics)" title="Morphology (linguistics)">morphological analysis</a>) or <a href="Paragraph" title="Paragraph">paragraphs</a>.
</p>
<div class="mw-heading mw-heading2"><h2 id="Automatic_segmentation_approaches">Automatic segmentation approaches</h2></div>
<p>Automatic segmentation is the problem in <a href="Natural_language_processing" title="Natural language processing">natural language processing</a> of implementing a computer process to segment text.
</p><p>When punctuation and similar clues are not consistently available, the segmentation task often requires fairly non-trivial techniques, such as statistical decision-making, large dictionaries, as well as consideration of syntactic and semantic constraints. Effective natural language processing systems and text segmentation tools usually operate on text in specific domains and sources. As an example, processing text used in medical records is a very different problem than processing news articles or real estate advertisements.
</p><p>The process of developing text segmentation tools starts with collecting a large corpus of text in an application domain. There are two general approaches:
</p>
<ul><li>Manual analysis of text and writing custom software</li>
<li>Annotate the sample corpus with boundary information and use <a href="Machine_learning" title="Machine learning">machine learning</a></li></ul>
<p>Some text segmentation systems take advantage of any markup like HTML and know document formats like PDF to provide additional evidence for sentence and paragraph boundaries.
</p>
<div class="mw-heading mw-heading2"><h2 id="See_also">See also</h2></div>
<ul><li><a href="Syllabification" title="Syllabification">Hyphenation</a></li>
<li><a href="Natural_language_processing" title="Natural language processing">Natural language processing</a></li>
<li><a href="Speech_segmentation" title="Speech segmentation">Speech segmentation</a></li>
<li><a href="Lexical_analysis" title="Lexical analysis">Lexical analysis</a></li>
<li><a href="Word_count" title="Word count">Word count</a></li>
<li><a href="Line_wrap_and_word_wrap" class="mw-redirect" title="Line wrap and word wrap">Line breaking</a></li>
<li><a href="Image_segmentation" title="Image segmentation">Image segmentation</a></li></ul>
<div class="mw-heading mw-heading2"><h2 id="References">References</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1239543626">
/* start https://en.wikipedia.org/ */
.mw-parser-output .reflist{margin-bottom:0.5em;list-style-type:decimal}@media screen{.mw-parser-output .reflist{font-size:90%}}.mw-parser-output .reflist .references{font-size:100%;margin-bottom:0;list-style-type:inherit}.mw-parser-output .reflist-columns-2{column-width:30em}.mw-parser-output .reflist-columns-3{column-width:25em}.mw-parser-output .reflist-columns{margin-top:0.3em}.mw-parser-output .reflist-columns ol{margin-top:0}.mw-parser-output .reflist-columns li{page-break-inside:avoid;break-inside:avoid-column}.mw-parser-output .reflist-upper-alpha{list-style-type:upper-alpha}.mw-parser-output .reflist-upper-roman{list-style-type:upper-roman}.mw-parser-output .reflist-lower-alpha{list-style-type:lower-alpha}.mw-parser-output .reflist-lower-greek{list-style-type:lower-greek}.mw-parser-output .reflist-lower-roman{list-style-type:lower-roman}
/* end https://en.wikipedia.org/ */
</style><div class="reflist">
<div class="mw-references-wrap"><ol class="references">
<li id="cite_note-1"><span class="mw-cite-backlink"><b><a href="#cite_ref-1">^</a></b></span> <span class="reference-text"><a rel="nofollow" class="external text" href="http://unicode.org/reports/tr29/">UAX #29</a></span>
</li>
<li id="cite_note-2"><span class="mw-cite-backlink"><b><a href="#cite_ref-2">^</a></b></span> <span class="reference-text"><style data-mw-deduplicate="TemplateStyles:r1238218222">
/* start https://en.wikipedia.org/ */
.mw-parser-output cite.citation{font-style:inherit;word-wrap:break-word}.mw-parser-output .citation q{quotes:"\"""\"""'""'"}.mw-parser-output .citation:target{background-color:rgba(0,127,255,0.133)}.mw-parser-output .id-lock-free.id-lock-free a{background:url("./mw/Lock-green.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-limited.id-lock-limited a,.mw-parser-output .id-lock-registration.id-lock-registration a{background:url("./mw/Lock-gray-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-subscription.id-lock-subscription a{background:url("./mw/Lock-red-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .cs1-ws-icon a{background:url("./mw/Wikisource-logo.svg")right 0.1em center/12px no-repeat}body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-free a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-limited a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-registration a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-subscription a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .cs1-ws-icon a{background-size:contain;padding:0 1em 0 0}.mw-parser-output .cs1-code{color:inherit;background:inherit;border:none;padding:inherit}.mw-parser-output .cs1-hidden-error{display:none;color:var(--color-error,#d33)}.mw-parser-output .cs1-visible-error{color:var(--color-error,#d33)}.mw-parser-output .cs1-maint{display:none;color:#085;margin-left:0.3em}.mw-parser-output .cs1-kern-left{padding-left:0.2em}.mw-parser-output .cs1-kern-right{padding-right:0.2em}.mw-parser-output .citation .mw-selflink{font-weight:inherit}@media screen{.mw-parser-output .cs1-format{font-size:95%}html.skin-theme-clientpref-night .mw-parser-output .cs1-maint{color:#18911f}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .cs1-maint{color:#18911f}}
/* end https://en.wikipedia.org/ */
</style><cite id="CITEREFZhang1998" class="citation journal cs1 cs1-prop-script cs1-prop-foreign-lang-source">Zhang, Xiao-heng (1998). <a rel="nofollow" class="external text" href="http://jcip.cipsc.org.cn/CN/Y1998/V12/I3/58"><bdi lang="zh">也谈汉语书面语的分词问题——分词连写十大好处</bdi></a> [Written Chinese Word-Segmentation Revisited: Ten advantages of word-segmented writing]. <i>中文信息学报</i> <bdi lang="zh">中文信息学报</bdi> [<i><a href="Journal_of_Chinese_Information_Processing" title="Journal of Chinese Information Processing">Journal of Chinese Information Processing</a></i>] (in Simplified Chinese). <b>12</b> (3): <span class="nowrap">58–</span>64<span class="reference-accessdate">. Retrieved <span class="nowrap">31 March</span> 2025</span>.</cite></span>
</li>
<li id="cite_note-3"><span class="mw-cite-backlink"><b><a href="#cite_ref-3">^</a></b></span> <span class="reference-text"><cite id="CITEREFChoi2000" class="citation conference cs1">Choi, Freddy Y. Y. (2000). <a rel="nofollow" class="external text" href="https://aclanthology.org/A00-2004/">"Advances in domain independent linear text segmentation"</a>. <i>Proceedings of the 1st Meeting of the North American Chapter of the Association for Computational Linguistics (ANLP-NAACL-00)</i>. pp. <span class="nowrap">26–</span>33. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/cs/0003083">cs/0003083</a></span><span class="reference-accessdate">. Retrieved <span class="nowrap">31 March</span> 2025</span>.</cite></span>
</li>
<li id="cite_note-4"><span class="mw-cite-backlink"><b><a href="#cite_ref-4">^</a></b></span> <span class="reference-text"><cite id="CITEREFReynar1998" class="citation thesis cs1">Reynar, Jeffrey C. (1998). <a rel="nofollow" class="external text" href="https://repository.upenn.edu/handle/20.500.14332/37673"><i>Topic Segmentation: Algorithms and Applications</i></a> <span class="cs1-format">(PDF)</span> (PhD thesis). <a href="University_of_Pennsylvania" title="University of Pennsylvania">University of Pennsylvania</a>. IRCS-98-21<span class="reference-accessdate">. Retrieved <span class="nowrap">31 March</span> 2025</span>.</cite></span>
</li>
</ol></div></div>
<p><br>
</p>
<div class="navbox-styles"><style data-mw-deduplicate="TemplateStyles:r1129693374">
/* start https://en.wikipedia.org/ */
.mw-parser-output .hlist dl,.mw-parser-output .hlist ol,.mw-parser-output .hlist ul{margin:0;padding:0}.mw-parser-output .hlist dd,.mw-parser-output .hlist dt,.mw-parser-output .hlist li{margin:0;display:inline}.mw-parser-output .hlist.inline,.mw-parser-output .hlist.inline dl,.mw-parser-output .hlist.inline ol,.mw-parser-output .hlist.inline ul,.mw-parser-output .hlist dl dl,.mw-parser-output .hlist dl ol,.mw-parser-output .hlist dl ul,.mw-parser-output .hlist ol dl,.mw-parser-output .hlist ol ol,.mw-parser-output .hlist ol ul,.mw-parser-output .hlist ul dl,.mw-parser-output .hlist ul ol,.mw-parser-output .hlist ul ul{display:inline}.mw-parser-output .hlist .mw-empty-li{display:none}.mw-parser-output .hlist dt::after{content:": "}.mw-parser-output .hlist dd::after,.mw-parser-output .hlist li::after{content:" · ";font-weight:bold}.mw-parser-output .hlist dd:last-child::after,.mw-parser-output .hlist dt:last-child::after,.mw-parser-output .hlist li:last-child::after{content:none}.mw-parser-output .hlist dd dd:first-child::before,.mw-parser-output .hlist dd dt:first-child::before,.mw-parser-output .hlist dd li:first-child::before,.mw-parser-output .hlist dt dd:first-child::before,.mw-parser-output .hlist dt dt:first-child::before,.mw-parser-output .hlist dt li:first-child::before,.mw-parser-output .hlist li dd:first-child::before,.mw-parser-output .hlist li dt:first-child::before,.mw-parser-output .hlist li li:first-child::before{content:" (";font-weight:normal}.mw-parser-output .hlist dd dd:last-child::after,.mw-parser-output .hlist dd dt:last-child::after,.mw-parser-output .hlist dd li:last-child::after,.mw-parser-output .hlist dt dd:last-child::after,.mw-parser-output .hlist dt dt:last-child::after,.mw-parser-output .hlist dt li:last-child::after,.mw-parser-output .hlist li dd:last-child::after,.mw-parser-output .hlist li dt:last-child::after,.mw-parser-output .hlist li li:last-child::after{content:")";font-weight:normal}.mw-parser-output .hlist ol{counter-reset:listitem}.mw-parser-output .hlist ol>li{counter-increment:listitem}.mw-parser-output .hlist ol>li::before{content:" "counter(listitem)"\a0 "}.mw-parser-output .hlist dd ol>li:first-child::before,.mw-parser-output .hlist dt ol>li:first-child::before,.mw-parser-output .hlist li ol>li:first-child::before{content:" ("counter(listitem)"\a0 "}
/* end https://en.wikipedia.org/ */
</style><style data-mw-deduplicate="TemplateStyles:r1236075235">
/* start https://en.wikipedia.org/ */
.mw-parser-output .navbox{box-sizing:border-box;border:1px solid #a2a9b1;width:100%;clear:both;font-size:88%;text-align:center;padding:1px;margin:1em auto 0}.mw-parser-output .navbox .navbox{margin-top:0}.mw-parser-output .navbox+.navbox,.mw-parser-output .navbox+.navbox-styles+.navbox{margin-top:-1px}.mw-parser-output .navbox-inner,.mw-parser-output .navbox-subgroup{width:100%}.mw-parser-output .navbox-group,.mw-parser-output .navbox-title,.mw-parser-output .navbox-abovebelow{padding:0.25em 1em;line-height:1.5em;text-align:center}.mw-parser-output .navbox-group{white-space:nowrap;text-align:right}.mw-parser-output .navbox,.mw-parser-output .navbox-subgroup{background-color:#fdfdfd}.mw-parser-output .navbox-list{line-height:1.5em;border-color:#fdfdfd}.mw-parser-output .navbox-list-with-group{text-align:left;border-left-width:2px;border-left-style:solid}.mw-parser-output tr+tr>.navbox-abovebelow,.mw-parser-output tr+tr>.navbox-group,.mw-parser-output tr+tr>.navbox-image,.mw-parser-output tr+tr>.navbox-list{border-top:2px solid #fdfdfd}.mw-parser-output .navbox-title{background-color:#ccf}.mw-parser-output .navbox-abovebelow,.mw-parser-output .navbox-group,.mw-parser-output .navbox-subgroup .navbox-title{background-color:#ddf}.mw-parser-output .navbox-subgroup .navbox-group,.mw-parser-output .navbox-subgroup .navbox-abovebelow{background-color:#e6e6ff}.mw-parser-output .navbox-even{background-color:#f7f7f7}.mw-parser-output .navbox-odd{background-color:transparent}.mw-parser-output .navbox .hlist td dl,.mw-parser-output .navbox .hlist td ol,.mw-parser-output .navbox .hlist td ul,.mw-parser-output .navbox td.hlist dl,.mw-parser-output .navbox td.hlist ol,.mw-parser-output .navbox td.hlist ul{padding:0.125em 0}.mw-parser-output .navbox .navbar{display:block;font-size:100%}.mw-parser-output .navbox-title .navbar{float:left;text-align:left;margin-right:0.5em}body.skin--responsive .mw-parser-output .navbox-image img{max-width:none!important}@media print{body.ns-0 .mw-parser-output .navbox{display:none!important}}
/* end https://en.wikipedia.org/ */
</style></div><div role="navigation" class="navbox" aria-labelledby="Natural_language_processing454" style="padding:3px"><table class="nowraplinks hlist mw-collapsible autocollapse navbox-inner" style="border-spacing:0;background:transparent;color:inherit"><tbody><tr><th scope="col" class="navbox-title" colspan="2"><style data-mw-deduplicate="TemplateStyles:r1239400231">
/* start https://en.wikipedia.org/ */
.mw-parser-output .navbar{display:inline;font-size:88%;font-weight:normal}.mw-parser-output .navbar-collapse{float:left;text-align:left}.mw-parser-output .navbar-boxtext{word-spacing:0}.mw-parser-output .navbar ul{display:inline-block;white-space:nowrap;line-height:inherit}.mw-parser-output .navbar-brackets::before{margin-right:-0.125em;content:"[ "}.mw-parser-output .navbar-brackets::after{margin-left:-0.125em;content:" ]"}.mw-parser-output .navbar li{word-spacing:-0.125em}.mw-parser-output .navbar a>span,.mw-parser-output .navbar a>abbr{text-decoration:inherit}.mw-parser-output .navbar-mini abbr{font-variant:small-caps;border-bottom:none;text-decoration:none;cursor:inherit}.mw-parser-output .navbar-ct-full{font-size:114%;margin:0 7em}.mw-parser-output .navbar-ct-mini{font-size:114%;margin:0 4em}html.skin-theme-clientpref-night .mw-parser-output .navbar li a abbr{color:var(--color-base)!important}@media(prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .navbar li a abbr{color:var(--color-base)!important}}@media print{.mw-parser-output .navbar{display:none!important}}
/* end https://en.wikipedia.org/ */
</style><div id="Natural_language_processing454" style="font-size:114%;margin:0 4em"><a href="Natural_language_processing" title="Natural language processing">Natural language processing</a></div></th></tr><tr><th scope="row" class="navbox-group" style="width:1%">General terms</th><td class="navbox-list-with-group navbox-list navbox-odd" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="AI-complete" title="AI-complete">AI-complete</a></li>
<li><a href="Bag-of-words_model" title="Bag-of-words model">Bag-of-words</a></li>
<li><a href="N-gram" title="N-gram"><i>n</i>-gram</a>
<ul><li><a href="Bigram" title="Bigram">Bigram</a></li>
<li><a href="Trigram" title="Trigram">Trigram</a></li></ul></li>
<li><a href="Computational_linguistics" title="Computational linguistics">Computational linguistics</a></li>
<li><a href="Natural_language_understanding" title="Natural language understanding">Natural language understanding</a></li>
<li><a href="Stop_word" title="Stop word">Stop words</a></li>
<li><a href="Text_processing" title="Text processing">Text processing</a></li></ul>
</div></td></tr><tr><th scope="row" class="navbox-group" style="width:1%"><a href="Text_mining" title="Text mining">Text analysis</a></th><td class="navbox-list-with-group navbox-list navbox-even" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="Argument_mining" title="Argument mining">Argument mining</a></li>
<li><a href="Collocation_extraction" title="Collocation extraction">Collocation extraction</a></li>
<li><a href="Concept_mining" title="Concept mining">Concept mining</a></li>
<li><a href="Coreference#Coreference_resolution" title="Coreference">Coreference resolution</a></li>
<li><a href="Deep_linguistic_processing" title="Deep linguistic processing">Deep linguistic processing</a></li>
<li><a href="Distant_reading" title="Distant reading">Distant reading</a></li>
<li><a href="Information_extraction" title="Information extraction">Information extraction</a></li>
<li><a href="Named-entity_recognition" title="Named-entity recognition">Named-entity recognition</a></li>
<li><a href="Ontology_learning" title="Ontology learning">Ontology learning</a></li>
<li><a href="Parsing" title="Parsing">Parsing</a>
<ul><li><a href="Semantic_parsing" title="Semantic parsing">Semantic parsing</a></li>
<li><a href="Syntactic_parsing_(computational_linguistics)" title="Syntactic parsing (computational linguistics)">Syntactic parsing</a></li></ul></li>
<li><a href="Part-of-speech_tagging" title="Part-of-speech tagging">Part-of-speech tagging</a></li>
<li><a href="Semantic_analysis_(machine_learning)" title="Semantic analysis (machine learning)">Semantic analysis</a></li>
<li><a href="Semantic_role_labeling" title="Semantic role labeling">Semantic role labeling</a></li>
<li><a href="Semantic_decomposition_(natural_language_processing)" title="Semantic decomposition (natural language processing)">Semantic decomposition</a></li>
<li><a href="Semantic_similarity" title="Semantic similarity">Semantic similarity</a></li>
<li><a href="Sentiment_analysis" title="Sentiment analysis">Sentiment analysis</a></li></ul>
<ul><li><a href="Terminology_extraction" title="Terminology extraction">Terminology extraction</a></li>
<li><a href="Text_mining" title="Text mining">Text mining</a></li>
<li><a href="Textual_entailment" title="Textual entailment">Textual entailment</a></li>
<li><a href="Truecasing" title="Truecasing">Truecasing</a></li>
<li><a href="Word-sense_disambiguation" title="Word-sense disambiguation">Word-sense disambiguation</a></li>
<li><a href="Word-sense_induction" title="Word-sense induction">Word-sense induction</a></li></ul>
</div><table class="nowraplinks navbox-subgroup" style="border-spacing:0"><tbody><tr><th id="Text_segmentation21" scope="row" class="navbox-group" style="width:1%"></th><td class="navbox-list-with-group navbox-list navbox-odd" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="Compound-term_processing" title="Compound-term processing">Compound-term processing</a></li>
<li><a href="Lemmatisation" class="mw-redirect" title="Lemmatisation">Lemmatisation</a></li>
<li><a href="Lexical_analysis" title="Lexical analysis">Lexical analysis</a></li>
<li><a href="Shallow_parsing" title="Shallow parsing">Text chunking</a></li>
<li><a href="Stemming" title="Stemming">Stemming</a></li>
<li><a href="Sentence_boundary_disambiguation" title="Sentence boundary disambiguation">Sentence segmentation</a></li>
<li><a href="Word#Word_boundaries" title="Word">Word segmentation</a></li></ul>
</div></td></tr></tbody></table><div>
</div></td></tr><tr><th scope="row" class="navbox-group" style="width:1%"><a href="Automatic_summarization" title="Automatic summarization">Automatic summarization</a></th><td class="navbox-list-with-group navbox-list navbox-even" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="Multi-document_summarization" title="Multi-document summarization">Multi-document summarization</a></li>
<li><a href="Sentence_extraction" title="Sentence extraction">Sentence extraction</a></li>
<li><a href="Text_simplification" title="Text simplification">Text simplification</a></li></ul>
</div></td></tr><tr><th scope="row" class="navbox-group" style="width:1%"><a href="Machine_translation" title="Machine translation">Machine translation</a></th><td class="navbox-list-with-group navbox-list navbox-odd" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="Computer-assisted_translation" title="Computer-assisted translation">Computer-assisted</a></li>
<li><a href="Example-based_machine_translation" title="Example-based machine translation">Example-based</a></li>
<li><a href="Rule-based_machine_translation" title="Rule-based machine translation">Rule-based</a></li>
<li><a href="Statistical_machine_translation" title="Statistical machine translation">Statistical</a></li>
<li><a href="Transfer-based_machine_translation" title="Transfer-based machine translation">Transfer-based</a></li>
<li><a href="Neural_machine_translation" title="Neural machine translation">Neural</a></li></ul>
</div></td></tr><tr><th scope="row" class="navbox-group" style="width:1%"><a href="Distributional_semantics" title="Distributional semantics">Distributional semantics</a> models</th><td class="navbox-list-with-group navbox-list navbox-even" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="BERT_(language_model)" title="BERT (language model)">BERT</a></li>
<li><a href="Document-term_matrix" title="Document-term matrix">Document-term matrix</a></li>
<li><a href="Explicit_semantic_analysis" title="Explicit semantic analysis">Explicit semantic analysis</a></li>
<li><a href="FastText" title="FastText">fastText</a></li>
<li><a href="GloVe" title="GloVe">GloVe</a></li>
<li><a href="Language_model" title="Language model">Language model</a> (<a href="Large_language_model" title="Large language model">large</a>)</li>
<li><a href="Latent_semantic_analysis" title="Latent semantic analysis">Latent semantic analysis</a></li>
<li><a href="Seq2seq" title="Seq2seq">Seq2seq</a></li>
<li><a href="Word_embedding" title="Word embedding">Word embedding</a></li>
<li><a href="Word2vec" title="Word2vec">Word2vec</a></li></ul>
</div></td></tr><tr><th scope="row" class="navbox-group" style="width:1%"><a href="Language_resource" title="Language resource">Language resources</a>,<br>datasets and corpora</th><td class="navbox-list-with-group navbox-list navbox-odd" style="width:100%;padding:0"><div style="padding:0 0.25em"></div><table class="nowraplinks navbox-subgroup" style="border-spacing:0"><tbody><tr><th scope="row" class="navbox-group" style="width:1%">Types and<br>standards</th><td class="navbox-list-with-group navbox-list navbox-odd" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="Corpus_linguistics" title="Corpus linguistics">Corpus linguistics</a></li>
<li><a href="Lexical_resource" title="Lexical resource">Lexical resource</a></li>
<li><a href="Linguistic_Linked_Open_Data" title="Linguistic Linked Open Data">Linguistic Linked Open Data</a></li>
<li><a href="Machine-readable_dictionary" title="Machine-readable dictionary">Machine-readable dictionary</a></li>
<li><a href="Parallel_text" title="Parallel text">Parallel text</a></li>
<li><a href="PropBank" title="PropBank">PropBank</a></li>
<li><a href="Semantic_network" title="Semantic network">Semantic network</a></li>
<li><a href="Simple_Knowledge_Organization_System" title="Simple Knowledge Organization System">Simple Knowledge Organization System</a></li>
<li><a href="Speech_corpus" title="Speech corpus">Speech corpus</a></li>
<li><a href="Text_corpus" title="Text corpus">Text corpus</a></li>
<li><a href="Thesaurus_(information_retrieval)" title="Thesaurus (information retrieval)">Thesaurus (information retrieval)</a></li>
<li><a href="Treebank" title="Treebank">Treebank</a></li>
<li><a href="Universal_Dependencies" title="Universal Dependencies">Universal Dependencies</a></li></ul>
</div></td></tr><tr><th scope="row" class="navbox-group" style="width:1%">Data</th><td class="navbox-list-with-group navbox-list navbox-even" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="BabelNet" title="BabelNet">BabelNet</a></li>
<li><a href="Bank_of_English" title="Bank of English">Bank of English</a></li>
<li><a href="DBpedia" title="DBpedia">DBpedia</a></li>
<li><a href="FrameNet" title="FrameNet">FrameNet</a></li>
<li><a href="Google_Ngram_Viewer" class="mw-redirect" title="Google Ngram Viewer">Google Ngram Viewer</a></li>
<li><a href="UBY" title="UBY">UBY</a></li>
<li><a href="WordNet" title="WordNet">WordNet</a></li>
<li><a href="Wikidata" title="Wikidata">Wikidata</a></li></ul>
</div></td></tr></tbody></table><div></div></td></tr><tr><th scope="row" class="navbox-group" style="width:1%"><a href="Automatic_identification_and_data_capture" title="Automatic identification and data capture">Automatic identification<br>and data capture</a></th><td class="navbox-list-with-group navbox-list navbox-odd" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="Speech_recognition" title="Speech recognition">Speech recognition</a></li>
<li><a href="Speech_segmentation" title="Speech segmentation">Speech segmentation</a></li>
<li><a href="Speech_synthesis" title="Speech synthesis">Speech synthesis</a></li>
<li><a href="Natural_language_generation" title="Natural language generation">Natural language generation</a></li>
<li><a href="Optical_character_recognition" title="Optical character recognition">Optical character recognition</a></li></ul>
</div></td></tr><tr><th scope="row" class="navbox-group" style="width:1%"><a href="Topic_model" title="Topic model">Topic model</a></th><td class="navbox-list-with-group navbox-list navbox-even" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="Document_classification" title="Document classification">Document classification</a></li>
<li><a href="Latent_Dirichlet_allocation" title="Latent Dirichlet allocation">Latent Dirichlet allocation</a></li>
<li><a href="Pachinko_allocation" title="Pachinko allocation">Pachinko allocation</a></li></ul>
</div></td></tr><tr><th scope="row" class="navbox-group" style="width:1%"><a href="Computer-assisted_reviewing" title="Computer-assisted reviewing">Computer-assisted<br>reviewing</a></th><td class="navbox-list-with-group navbox-list navbox-odd" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="Automated_essay_scoring" title="Automated essay scoring">Automated essay scoring</a></li>
<li><a href="Concordancer" title="Concordancer">Concordancer</a></li>
<li><a href="Grammar_checker" title="Grammar checker">Grammar checker</a></li>
<li><a href="Predictive_text" title="Predictive text">Predictive text</a></li>
<li><a href="Pronunciation_assessment" title="Pronunciation assessment">Pronunciation assessment</a></li>
<li><a href="Spell_checker" title="Spell checker">Spell checker</a></li></ul>
</div></td></tr><tr><th scope="row" class="navbox-group" style="width:1%"><a href="Natural-language_user_interface" title="Natural-language user interface">Natural language<br>user interface</a></th><td class="navbox-list-with-group navbox-list navbox-even" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="Chatbot" title="Chatbot">Chatbot</a></li>
<li><a href="Interactive_fiction" title="Interactive fiction">Interactive fiction</a></li>
<li><a href="Question_answering" title="Question answering">Question answering</a></li>
<li><a href="Virtual_assistant" title="Virtual assistant">Virtual assistant</a></li>
<li><a href="Voice_user_interface" title="Voice user interface">Voice user interface</a></li></ul>
</div></td></tr><tr><th scope="row" class="navbox-group" style="width:1%">Related</th><td class="navbox-list-with-group navbox-list navbox-odd" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="Formal_semantics_(natural_language)" title="Formal semantics (natural language)">Formal semantics</a></li>
<li><a href="Hallucination_(artificial_intelligence)" title="Hallucination (artificial intelligence)">Hallucination</a></li>
<li><a href="Natural_Language_Toolkit" title="Natural Language Toolkit">Natural Language Toolkit</a></li>
<li><a href="SpaCy" title="SpaCy">spaCy</a></li></ul>
</div></td></tr></tbody></table></div></div><!--htdig_noindex--><div><div class="zim-footer">
This article is issued from <a class="external text" title="Last edited on 2025-04-30" href="https://en.wikipedia.org/wiki/?title=Text_segmentation&oldid=1288110012">Wikipedia</a>. The text is available under <a class="external text" href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">Creative Commons Attribution-Share Alike 4.0</a> unless otherwise noted. Additional terms may apply for the media files.
</div>
</div><!--/htdig_noindex--></div>
</div>
</main>
</div>
</div>
</div>
</body></html>